Papers with AI agents
Systematic Biases in LLM Simulations of Debates (2024.emnlp-main)
Copied to clipboard
| Challenge: | Current research suggests that LLM-based agents become increasingly human-like in their performance, sparking interest in using these AI agents as substitutes for human participants in behavioral studies. |
| Approach: | They propose to use LLMs to simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes. |
| Outcome: | The proposed model can simulate political debates on topics that are important aspects of people’s day-to-day lives and decision-making processes. |
Generating OpenAPI Specifications from Online API Documentation with Large Language Models (2025.acl-industry)
Copied to clipboard
Koren Lazar, Matan Vetzler, Kiran Kate, Jason Tsay, David Boaz, Himanshu Gupta, Avraham Shinnar, Rohith D Vallam, David Amid, Esther Goldbraich, Jim Laredo, Ateret Anaby Tavor
| Challenge: | API specifications are often presented as unstructured HTML pages, requiring external users to manually convert it into a structured format. |
| Approach: | They propose a framework that transforms long API documentation pages into consistent, machine-readable API specifications. |
| Outcome: | The proposed framework generalizes well across hundreds of APIs and produces valid OpenAPI specifications that encapsulate most of the information from the original documentation. |
The Price of Thought: A Multilingual Analysis of Reasoning, Performance, and Cost of Negotiation in Large Language Models (2026.findings-eacl)
Copied to clipboard
Sherzod Hakimov, Roland Bernard, Tim Leiber, Karl Osswald, Kristina Richert, Ruilin Yang, Raffaella Bernardi, David Schlangen
| Challenge: | Negotiation is a fundamental challenge for AI agents as it requires an ability to reason strategically, model opponents, and balance cooperation with competition. |
| Approach: | They propose to use a self-play setup to compare commercial and open-weight large language models to their vanilla counterparts in three different languages to examine trade-offs between performance and cost. |
| Outcome: | The proposed model improves GPT-5's performance by 31.4 % while increasing its cost by nearly 400 %. |
FOFO: A Benchmark to Evaluate LLMs’ Format-Following Capability (2024.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks fail to assess large language models’ format-following proficiency adequately. |
| Approach: | They propose a benchmark to evaluate large language models' ability to follow complex, domain-specific formats. |
| Outcome: | The proposed framework evaluates large language models' ability to follow complex, domain-specific formats across open-source and closed-source models. |
A11y-Compressor: A Framework for Enhancing the Efficiency of GUI Agent Observations through Visual Context Reconstruction and Redundancy Reduction (2026.acl-srw)
Copied to clipboard
| Challenge: | Existing approaches to grounding GUI environments are categorized into image-based and text-based representations. |
| Approach: | They propose a framework that transforms linearized accessibility trees into compact and structured representations. |
| Outcome: | The proposed framework reduces input tokens to 22% of the original while improving task success rates by 5.1 percentage points on average. |
Don’t Trust Generative Agents to Mimic Communication on Social Networks Unless You Benchmarked their Empirical Realism (2026.eacl-long)
Copied to clipboard
| Challenge: | Social media platforms face mounting regulatory pressure worldwide . obtaining evidence regarding platform risks remains challenging . |
| Approach: | They propose a formal framework for simulation of social networks before focusing on imitating user communication. |
| Outcome: | The proposed model can replicate human behavior with sufficient realism to perform the task. |
Interactive Training: Feedback-Driven Neural Network Optimization (2025.emnlp-demos)
Copied to clipboard
| Challenge: | In traditional neural network training, static optimization methods lack flexibility and responsiveness . authors demonstrate that Interactive Training provides superior training stability and reduced sensitivity to initial hyperparameters . |
| Approach: | They propose an open-source framework that enables real-time feedback-driven optimization of neural networks by human experts or automated AI agents. |
| Outcome: | The proposed framework achieves superior training stability, reduced sensitivity to initial hyperparameters, and improved adaptability to evolving user needs. |
SlackAgents: Scalable Collaboration of AI Agents in Workspaces (2025.emnlp-demos)
Copied to clipboard
Zhiwei Liu, Weiran Yao, Zuxin Liu, Juntao Tan, Jianguo Zhang, Frank Wang, Sukhandeep Nahal, Huan Wang, Shelby Heinecke, Silvio Savarese, Caiming Xiong
| Challenge: | Existing open-source frameworks like LangChain and LlamaIndex fail to integrate into daily workflows, resulting in limited daily usage for work. |
| Approach: | They propose a multi-agent library for scalable management and collaboration of AI agents on Slack. |
| Outcome: | The proposed framework offers instant AI integration into organizational workflows and facilitates scalable collaboration, allowing for effective communication and task orchestration. |
Grounding Open-Domain Instructions to Automate Web Support Tasks (2021.naacl-main)
Copied to clipboard
| Challenge: | RUSS is a task and dataset to ground natural language instructions on the web to perform previously unseen tasks. |
| Approach: | They build a task and dataset to ground AI agents from open-domain, step-by-step instructions on the web. |
| Outcome: | The proposed model outperforms existing models that map instructions to actions without WebLang. |
MindCraft: Theory of Mind Modeling for Situated Dialogue in Collaborative Tasks (2021.emnlp-main)
Copied to clipboard
| Challenge: | Creating embodied, situated agents able to move in, communicate naturally about, and collaborate on human terms in the physical world has been a persisting goal in artificial intelligence (Winograd, 1972). |
| Approach: | They propose to use a 3D Minecraft dataset to model the beliefs of human partners in situ to enable theory of mind modeling in situated interactions. |
| Outcome: | The proposed model can be used to model human collaborative behaviors in the 3D virtual blocks world of Minecraft. |
Improving Grounded Language Understanding in a Collaborative Environment by Interacting with Agents Through Help Feedback (2024.findings-eacl)
Copied to clipboard
| Challenge: | In many approaches to Natural Language Processing tasks, language is inherently interactive. |
| Approach: | They propose to use human-AI collaboration to improve human-human interaction by providing feedback that the agent can understand and utilize. |
| Outcome: | The proposed task is an interactive grounded language understanding task in a MineCraft-like world. |
AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents (2025.findings-acl)
Copied to clipboard
Yuxiang Chai, Siyuan Huang, Yazhe Niu, Han Xiao, Liang Liu, Guozhi Wang, Dingyu Zhang, Shuai Ren, Hongsheng Li
| Challenge: | a new dataset is being developed to improve the capabilities of mobile GUI-control agents. |
| Approach: | They propose a dataset designed for generalist mobile GUI-control agents . they use screenshots from popular mobile applications to create a detailed GUI-annotated dataset . |
| Outcome: | The Android Multi-annotation EXpo (AMEX) is a large-scale dataset for generalist mobile GUI-control agents . it includes screenshots from popular mobile applications, which are annotated at multiple levels . |
Towards LLM Agents for Earth Observation (2026.findings-acl)
Copied to clipboard
Chia Hsiang Kao, Wenting Zhao, Cheryl Lam, Aarush Umap, Shreelekha Revankar, Samuel Speas, Snehal Bhagat, Rajeev Datta, Cheng Perng Phoo, Utkarsh Mall, Carl Vondrick, Kavita Bala, Bharath Hariharan
| Challenge: | specialized automated systems for specific earth observation tasks lack flexibility for general-purpose, customized queries. |
| Approach: | They propose a coding benchmark of 408 yes/no questions from NASA Earth Observatory articles . they analyze the impact of using JavaScript API versus Python and the effect of providing documentation . |
| Outcome: | The proposed frameworks reduce errors by 60%, but are only marginally above random chance. |
Make The Most of Prior Data: A Solution for Interactive Text Summarization with Preference Feedback (2022.findings-naacl)
Copied to clipboard
Duy-Hung Nguyen, Nguyen Viet Dung Nghiem, Bao-Sinh Nguyen, Dung Tien Tien Le, Shahab Sabahi, Minh-Tien Nguyen, Hung Le
| Challenge: | a framework to train summarization models with preference feedback is proposed . human-in-the-loop (HITL) allows humans to actively participate in supervising AI systems . |
| Approach: | They propose a framework to train summarization models with preference feedback interactively. |
| Outcome: | The proposed framework improves ROUGE scores and sample-efficiency in active, few-shot and online settings. |
Ruler: A Model-Agnostic Method to Control Generated Length for Large Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models struggle to meet user’s needs when required to generate responses of a specific length due to their inherent difficulty in accurately perceiving numerical constraints. |
| Approach: | They propose a Target Length Generation Task and propose RULER, a model-agnostic approach that controls generated length for large language models. |
| Outcome: | The proposed model-agnostic approach improves instruction-following ability of large language models under length-constrained instructions and can generate appropriate MLT when length constraints are not explicitly provided. |
CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments (2025.naacl-long)
Copied to clipboard
Kung-Hsiang Huang, Akshara Prabhakar, Sidharth Dhawan, Yixin Mao, Huan Wang, Silvio Savarese, Caiming Xiong, Philippe Laban, Chien-Sheng Wu
| Challenge: | Existing benchmarks for evaluating CRM agents on work-related tasks are limited due to data privacy concerns. |
| Approach: | They propose a benchmark to evaluate AI agents on real-world CRM tasks . they use 16 commonly used industrial objects with high interconnectivity to simulate real data distributions. |
| Outcome: | The new benchmark evaluates AI agents on real-world customer service tasks . it includes 16 commonly used industrial objects with high interconnectivity . the results highlight the need for enhanced agent capabilities in function-calling and rule-following . |
Learning to Mediate Disparities Towards Pragmatic Communication (2022.acl-long)
Copied to clipboard
| Challenge: | Recent work explores pragmatic reasoning based on Rational Speech Act (RSA) and Theory of Mind in communication (Zhu et al., 2021). |
| Approach: | They propose a framework where the speaker attempts to learn the speaker-listener disparity and adjust the speech accordingly by adding a light-weighted disparity adjustment layer into working memory on top of speaker’s long-term memory system. |
| Outcome: | The proposed framework can learn and adapt to different types of listeners by adding a light-weighted disparity adjustment layer into working memory on top of speaker’s long-term memory system. |
Autonomous Workflow for Multimodal Fine-Grained Training Assistants Towards Mixed Reality (2024.findings-acl)
Copied to clipboard
Jiahuan Pei, Irene Viola, Haochen Huang, Junxiao Wang, Moonisa Ahsan, Fanghua Ye, Jiang Yiming, Yao Sai, Di Wang, Zhumin Chen, Pengjie Ren, Pablo Cesar
| Challenge: | a fine-grained, comprehensive understanding of multimodal environments remains under-explored. |
| Approach: | They propose an automated workflow for integrating AI agents into extended reality (XR) they propose a cerebral language agent that integrates LLM with memory, planning, and interaction with XR tools and a vision-language agent . |
| Outcome: | The proposed workflow integrates AI agents seamlessly into extended reality (XR) applications for fine-grained training. |
ReachAgent: Enhancing Mobile Agent via Page Reaching and Operation (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing mobile AI agents focus on most task-relevant elements at each step, leading to local optimal solutions and ignoring the overall GUI flow. |
| Approach: | They propose a mobile AI agent that breaks tasks into page reaching and operation subtasks and a framework that focuses on improving its task-completion abilities. |
| Outcome: | The proposed framework improves IoU accuracy and text accuracy by 7.12% and 7.69% on step-level and 4.72% and 4.63% on task-level compared to the SOTA agent. |
The Stackelberg Speaker: Optimizing Persuasive Communication in Social Deduction Games (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches focus on information processing and strategy selection, overlooking the significance of persuasive communication in social deduction games. |
| Approach: | They propose a reinforcement learning framework that trains agents to optimize influential utterances for persuasive impact by formalizing turn-based dialogue as a Stackelberg competition . |
| Outcome: | The proposed framework outperforms baselines across four social deduction benchmarks and shows that it is effective in persuasive communication. |
EmpathicStories++: A Multimodal Dataset for Empathy Towards Personal Experiences (2024.findings-acl)
Copied to clipboard
Jocelyn Shen, Yubin Kim, Mohit Hulse, Wazeer Zulfikar, Sharifa Alghowinem, Cynthia Breazeal, Hae Park
| Challenge: | Existing datasets for empathy modeling are limited in the ways they are not captured in the wild. |
| Approach: | They propose a multimodal dataset for empathy during personal experience sharing that contains 53 hours of video, audio, and text data of 41 participants. |
| Outcome: | The EmpathicStories++ dataset contains 53 hours of video, audio, and text data of 41 participants sharing vulnerable experiences and reading empathically resonant stories with an AI agent. |
Coding Agents with Multimodal Browsing are Generalist Problem Solvers (2026.findings-eacl)
Copied to clipboard
| Challenge: | specialized AI agents with task-specific tools or architectures fail to generalize beyond their intended scope. |
| Approach: | They propose a single-agent system with a modest number of general tools . they propose to generalize across software engineering, deep research and web browsing . |
| Outcome: | The proposed system achieves superior or competitive performance over specialized agents on three benchmarks. |
AI, Take the Wheel: What Drives Delegation and Trust in Human–Computer Cooperative Question Answering? (2026.findings-acl)
Copied to clipboard
Maharshi Gor, Yoo Yeon Sung, Yu Hou, Eve Fleisig, Zhu Irene Ying, Tianyi Zhou, Jordan Lee Boyd-Graber
| Challenge: | Human-AI collaboration is already happening, both in proactive delegation and deliberative adoption settings. |
| Approach: | They study delegating a task to AI without seeing its output and evaluating AI suggestions to decide whether to adopt them how AI output shapes final decisions. |
| Outcome: | The proposed game pairs 23 experts with 16 AI agents, capturing 387 delegation and 1440 adoption decisions. |
Attacks by Content: Automated Fact-checking is an AI Security Issue (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing defenses focus on detecting hidden commands but are ineffective against content attacks. |
| Approach: | They propose to repurpose retrieval-augmented generation (RAG) as a cognitive self-defense tool for agents. |
| Outcome: | The proposed approach is analogous to an existing task, automated fact-checking, and could be used to defend agents against content attacks. |
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games (2024.findings-acl)
Copied to clipboard
| Challenge: | In this study, we explore the application of Large Language Models (LLMs) in Jubensha, a Chinese detective role-playing game and a novel area in Artificial Intelligence (AI) driven gaming. |
| Approach: | They propose to use large language models to foster AI agent development in Jubensha, a Chinese detective role-playing game. |
| Outcome: | The proposed framework enables AI agents to engage in Jubensha games autonomously. |
LEDGER: Scaling Agentic Document Editing with Dependency-aware Graph Retrieval (2026.findings-acl)
Copied to clipboard
| Challenge: | Document editing requires full-context awareness of dependencies, but processing entire documents for each edit incurs prohibitive token costs and latency. |
| Approach: | a framework that constructs lightweight dependency graphs captures semantic relationships and structural hierarchies across document elements is proposed for agentic document editing . a scaLing agentic agentic framework is based on a dependency graph framework that captures dependencies and refactors function dependencies. |
| Outcome: | a new framework achieves 76 consistency versus 56 baseline while reducing token usage by 85 . the framework is based on a framework that captures semantic relationships and structural hierarchies across document elements . it can be used to improve document consistency, but it also reduces token costs and latency . |
Peering Behind the Shield: Guardrail Identification in Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Identifying guardrails in conversational AI agents is critical for identifying malicious content . identifying guardrail components in black-box AI agents poses security challenges . |
| Approach: | They propose a method that leverages guard-specific adversarial prompts to detect guardrails in black-box AI agents. |
| Outcome: | The proposed method achieves perfect classification accuracy in multiple scenarios. |
MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding (2024.findings-emnlp)
Copied to clipboard
Qinzhuo Wu, Weikai Xu, Wei Liu, Tao Tan, Liujian Liujianfeng, Ang Li, Jian Luan, Bin Wang, Shuo Shang
| Challenge: | Recent mobile AI agents based on VLMs lack basic mobile capabilities due to their pre-trained nature. |
| Approach: | They propose a mobile AI agent based on VLMs that includes additional pre-training stages to enhance both intra- and inter-UI understanding. |
| Outcome: | The proposed model outperforms existing VLMs on the Chinese mobile dataset Mobile3M . |
Beyond Screenshots: Evaluating VLMs’ Understanding of UI Animations (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent studies of Vision Language Models (VLMs) for UI understanding have focused primarily on static screenshots, leaving it unclear how well these models handle dynamic UI animations. |
| Approach: | They evaluate UI animation models' ability to perceive animation effects and interpret animation meaning . they use motion, context, and perceptual cues to probe factors affecting VLM performance . |
| Outcome: | The proposed model can detect primitive motion, but its interpretation is inconsistent . the proposed model is based on 300 annotated UI animation videos . |
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations (2024.findings-emnlp)
Copied to clipboard
Abhinav Joshi, Shaswati Saha, Divyaksh Shukla, Sriram Vema, Harsh Jhamtani, Manas Gaur, Ashutosh Modi
| Challenge: | Large Language Models (LLMs) have shown to be a great success in a wide range of applications ranging from regular NLP-based use cases to AI agents. |
| Approach: | They examine the robustness of existing MUL techniques for their ability to enable leakage-proof forgetting in LLMs. |
| Outcome: | The proposed methods can be used to enable leakage-proof forgetting in LLMs. |
Evolving Agents (2026.acl-long)
Copied to clipboard
| Challenge: | Current models are static entities incapable of compressing complexity of real world into generalisable concepts . authors: lack of endogenous mechanism for representation updating renders models vulnerable to domain mismatch and catastrophic forgetting . |
| Approach: | a meta-control system distils on-the-fly abstract representations of states, actions, goals . authors propose a paradigm for autonomous learning driven by pseudo-symbolic abstraction . |
| Outcome: | a meta-control system distils on-the-fly abstract representations of states, actions, goals . a novel approach resolves the domain mismatch problem and lays the groundwork for truly autonomous AI models . |
Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics (2025.emnlp-main)
Copied to clipboard
Jiarui Liu, Yueqi Song, Yunze Xiao, Mingqian Zheng, Lindia Tjuatja, Jana Schaich Borg, Mona T. Diab, Maarten Sap
| Challenge: | a study of multi-dimensional persona effects in AI-AI debates shows that personas influence moral stances and debate outcomes . political ideology and personality traits exert the strongest influence, according to our study . |
| Approach: | They propose to use a 6-dimensional persona space to simulate structured debates . they find political ideology and personality traits exert the strongest influence . |
| Outcome: | The study shows that personas affect moral stances and debate outcomes . political ideology and personality traits exert the strongest influence . |
Minimal Yet Big Impact: How AI Agent Back-channeling Enhances Conversational Engagement through Conversation Persistence and Context Richness (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Increasing use of AI agents in conversational services highlights the importance of back-channeling (BC) as an active listening strategy to enhance conversational engagement. |
| Approach: | They conducted an experiment with 55 participants to evaluate conversational engagement using both quantitative and qualitative metrics. |
| Outcome: | The results show that the Todak_BC and TodAK_NoBC groups have significantly higher conversational engagement than the Todask_NoB. |
Visual Inception: Compromising Long-term Planning in Agentic Recommenders via Multimodal Memory Poisoning (2026.acl-long)
Copied to clipboard
| Challenge: | Existing research focuses on prompt injection or immediate adversarial misclassification of user-uploaded images. |
| Approach: | They propose a dual-process defense framework inspired by human cognition to mitigate this vulnerability by injecting triggers into user-uploaded images that act as "sleeper agents" |
| Outcome: | The proposed framework achieves about 85% Goal-Hit Rate (GHR) while reducing the risk to 10% with configurable latency trade-offs. |
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent studies show that LLM-based agents struggle to perform in zero-shot scenarios. |
| Approach: | They propose a framework to quantify the behavior gap between AI agents and human experts . they propose to examine discrepancies in dialog acts, tool usage, and knowledge utilization . |
| Outcome: | The proposed framework measures the behavior gap between AI agents and human experts on task-oriented dialogs. |
REPRO-Bench: Can Agentic AI Systems Assess the Reproducibility of Social Science Research? (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for reproducing social science papers focus on reproducing results using provided code and data without assessing their consistency with the paper. |
| Approach: | They propose a benchmark to evaluate agentic AI systems' ability to automate reproducibility assessment. |
| Outcome: | The proposed benchmark oversimplifies real-world scenarios and lacks diversity in data formats and programming languages. |
Memory OS of AI Agent (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) face a shortage of long-term memory capabilities and limited personalization due to fixed context windows. |
| Approach: | They propose a Memory Operating System to achieve efficient memory management for AI agents . MemoryOS enables hierarchical memory integration and dynamic updating . |
| Outcome: | The proposed architecture enables hierarchical memory integration and dynamic updating. |
PAC-BENCH: Evaluating Multi-Agent Collaboration under Privacy Constraints (2026.findings-acl)
Copied to clipboard
Minjun Park, Donghyun Kim, Hyeonjong Ju, Seungwon Lim, Dongwook Choi, Taeyoon Kwon, Minju Kim, Jinyoung Yeo
| Challenge: | Recent research explores multi-agent systems where agents collaborate toward shared goals to handle complex tasks. |
| Approach: | They propose a benchmark for systematic evaluation of multi-agent collaboration under privacy constraints. |
| Outcome: | The proposed benchmark shows that privacy constraints degrade collaboration performance and make outcomes depend more on the initiating agent than the partner. |
Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents (2026.acl-long)
Copied to clipboard
| Challenge: | Small shifts in user behavior can cause sharp drops in agent performance . prior work has shown that LLMs lack robustness to real-world noise and small input perturbations. |
| Approach: | They propose a model-agnostic method for systematically stress testing AI agents that learns directions in activation space corresponding to steerable user traits. |
| Outcome: | The proposed method can be used to stress test AI agents in airline, retail, telecom, and telehealth domains. |